Papers with systematic methods
Thesis Proposal: Multimodal Benchmark for Music Understanding in Large Language Models (2026.eacl-srw)
Copied to clipboard
| Challenge: | Existing music-focused benchmarks are fragmented, largely single-modality, Western-centric . existing methods for evaluating MLLMs are lacking reproducibility and reliability . |
| Approach: | They propose to develop a musically multimodal benchmark that will integrate music into the benchmark. |
| Outcome: | The proposed benchmark will integrate culturally diverse musical material beyond the dominant Western canon. |
FINEST: Improving LLM Responses to Sensitive Topics Through Fine-Grained Evaluation (2026.findings-eacl)
Copied to clipboard
| Challenge: | Existing evaluation frameworks lack systematic methods to identify weaknesses in LLMs . Existing methods to evaluate LLM responses to sensitive topics are lacking . |
| Approach: | They propose a FINE-grained response evaluation taxonomy for sensitive topics that breaks down helpfulness and harmlessness into errors across three main categories: Content, Logic, and Appropriateness. |
| Outcome: | The proposed model outperforms refinement without guidance on Korean-sensitive questions . FINEST significantly improves the model responses across all three categories . |
Just Fine-tune Twice: Selective Differential Privacy for Large Language Models (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing approaches to protect language models from privacy leakage suffer from limited user control and low utility . et al., 2018: a novel framework that achieves SDP for state-of-the-art large transformer-based models. |
| Approach: | They propose a framework that applies differential privacy to large language models . they use redacted in-domain data to fine-tune the model with original in- domain data . |
| Outcome: | The proposed framework achieves strong utility compared to baselines. |
What does a Text Classifier Learn about Morality? An Explainable Method for Cross-Domain Comparison of Moral Rhetoric (2023.acl-long)
Copied to clipboard
Enrico Liscio, Oscar Araque, Lorenzo Gatti, Ionut Constantinescu, Catholijn Jonker, Kyriaki Kalimeri, Pradeep Kumar Murukannaiah
| Challenge: | Existing methods to analyze whether a text classifier learns the domain-specific expression of moral language are lacking. |
| Approach: | They propose a method to compare a supervised classifier’s representation of moral rhetoric across domains by exploring similarities and differences between moral concepts and domains. |
| Outcome: | The proposed method compares a supervised classifier’s representation of moral rhetoric across domains and domains. |